Human Genetics
○ Springer Science and Business Media LLC
Preprints posted in the last 90 days, ranked by how well they match Human Genetics's content profile, based on 28 papers previously published here. The average preprint has a 0.02% match score for this journal, so anything above that is already an above-average fit.
Uria-Regojo, G.; Fernandez-Caballero, L.; Lopez-Alcojor, A.; Lopez-Lopez, L.; Benitez, Y.; Rodilla, C.; Avila Fernandez, A.; Trujillo-Tiebas, M. J.; Osorio, A.; Corton, M.; Almoguera, B.; Ayuso, C.; Minguez, P.
Show abstract
Rare diseases (RDs) remain a major diagnostic challenge. Genetic and phenotypic heterogeneity, incomplete knowledge of disease mechanisms, and limitations in variant clinical interpretation leave many patients without a molecular diagnosis. Meanwhile, the growing volume of genomic data generated in clinical practice offers an opportunity to develop data-driven methodologies for exploring disease mechanisms and improving the reanalysis of unsolved cases. We aggregated real-world genomic data from 11,084 unrelated patients with suspected RD. Patients were clinically classified into 122 diseases. We built a multi-disease genomic variant frequency database (FJD-DB), which enabled the development of variant and gene-disease association scores by means of case-control subcohort comparisons across 32 disease groups. Functional enrichment analyses were then used to highlight disease-associated protein domains, pathways, biological processes, and phenotypes. Finally, the resulting knowledge was integrated into a data-driven framework for the guided reanalysis of unsolved RD patients applied to Inherited Retinal Dystrophies (IRD) patients as first use case. FJD-DB contained more than 45 million unique variants, including ~185,000 potentially pathogenic variants. Disease-specific analyses identified disease-associated pathogenic variants and highlighted both established and candidate disease genes. We detected 179 significantly enriched protein domains across 23 diseases, 124 Human Phenotype Ontology terms across 13 diseases, 79 Reactome pathways across 10 diseases, and 72 Gene Ontology biological processes across 8 diseases, revealing highly disease-specific functional signatures. Integration of disease-specific variant, gene, and functional association signals enabled the development of a data-driven framework for guided reanalysis of unsolved RD cases. Applied to more than 1,100 unsolved IRD cases, the framework generated clinically relevant findings in 26 patients, including four molecular diagnoses, seven candidate diagnoses, and 15 cases upgraded from non-informative findings to variants of uncertain significance. Aggregated real-world genomic data can be leveraged to identify disease-associated molecular signals generating novel biological hypotheses. A unified analytical framework provides a scalable strategy for knowledge discovery and guided reanalysis, facilitating the identification of overlooked and potentially novel genetic causes of RDs.
Rajueni, K.; Koskimaki, F.; Salo, V.; Pasanen, A.; Sliz, E.; Vanhala, S.; Reis, K.; Reigo, A.; FinnGen, ; Estonian Biobank Research Team, ; Palta, P.; Tasanen, K.; Liinamaa, J.; Kettunen, J.; Saarela, V.; Karjalainen, M. K.
Show abstract
Objective: The objective of this study was to detect genetic factors associated with dermatochalasis using a genome-wide association study (GWAS) across three large cohorts. Design: GWAS meta-analysis Participants: A total of 13,200 dermatochalasis cases and 962,513 controls were included. Methods: A GWAS meta-analysis of dermatochalasis combining data from the FinnGen, the Estonian Biobank and the UK Biobank was conducted. We also performed colocalization analyses, a phenome-wide association study and age-at-onset analysis, and assessed genetic correlations with various diseases and traits. Main outcome measures: Identification of genetic variants associated with dermatochalasis. Results: We identified 18 loci associated with dermatochalasis at genome-wide significance, 16 of which were novel. Most of these loci had genes involved in skin biology and cutaneous diseases, such as the genes encoding elastin (ELN) and Latent TGF-{beta} binding protein 1 (LTBP1). Phenome-wide association study revealed previous associations with morphology-related traits, while genetic correlation analysis highlighted multiple genetic correlations, especially with smoking and pain. Conclusions: We detected 18 genetic loci associated with dermatochalasis, characterized these loci in detail and demonstrated their relevance in skin biology and related processes. These findings give novel information on the genetic background of dermatochalasis and provide a solid basis for further research.
Jamalalail, B.; Khalifa, A.; Balan, B.; Bineshaq, S.; Advani, D.; Elsokary, H.; Dasuki, K.; Shiyas, S.; Soares, N. C.; Hanif, S.; Tharakan, S.; Mohamdi, Z.; Aburaidah, M.; Kuttiankandy, S.; Alsheikh-Ali, A.; Nassir, N.; El Bitar, M.; Uddin, M.
Show abstract
Consanguinity increases the risk of autosomal recessive disorders and may result in the co-segregation of multiple pathogenic variants within the same family. Although most affected families are explained by a single genetic diagnosis, multilocus pathogenic variation can produce complex and overlapping clinical phenotypes. We investigated a consanguineous Pakistani family with three affected siblings, including dizygotic twins presenting with neurodevelopmental disorder and hearing loss, using detailed clinical evaluation, long-read whole-genome sequencing, bulk transcriptomics, protein profiling, and segregation analysis to determine the underlying molecular diagnoses. One sibling presented with isolated non-syndromic hearing loss, whereas the dizygotic twins exhibited severe neurodevelopmental impairment characterized by global developmental delay, spastic quadriplegic cerebral palsy, microcephaly, and white matter abnormalities. Long-read whole-genome sequencing identified a homozygous start-loss variant in HPDL (c.3G>C) in both twins, consistent with HPDL-related neurodevelopmental disorder with progressive spasticity and brain white matter abnormalities (NEDSWMA). In addition, a novel homozygous nonsense variant in EPS8 (c.1294C>T) was identified in one twin and the sibling with isolated hearing loss, explaining the auditory phenotype. Long-read transcriptomic analysis demonstrated absence of detectable EPS8 transcripts in both individuals homozygous for the nonsense variant, providing transcript-level evidence consistent with a loss-of-function mechanism. Genome-wide comprehensive proteomic profiling (SomaScan) identified distinct protein abundance profiles across family members, with the most pronounced alterations observed in the twins affected by HPDL-related neurodevelopmental disease, particularly the individual harboring pathogenic variants in both EPS8 and HPDL. This study expands the mutational spectrum of EPS8 and highlights the independent segregation of two autosomal recessive disorders within the complex consanguineous family, resulting in distinct and blended phenotypes.
Tompson, S. W. J.; Graham, P.; Hadler, J.; Pasutto, F.; Whisenhunt, K. N.; Chakrabarti, S.; Young, T. L.; Craig, J. E.; Hewitt, A. W.; Siggs, O. M.; Hulleman, J. D.; Mackey, D. A.; Burdon, K. P.; Dubowsky, A.; Souzeau, E.
Show abstract
Pathogenic variants in the myocilin (MYOC) gene are the most common cause of Mendelian open-angle glaucoma. In 2022, the Clinical Genome Resource (ClinGen) Glaucoma Variant Curation Expert Panel (VCEP) published rule specifications for MYOC variant interpretation, including a pilot study of 81 variants. Here, we present the results of curating 271 MYOC variants reported in people with open-angle glaucoma using updated specification rules. Of all the variants, 11 were classified as benign (B), 45 as likely benign (LB), 166 as variants of uncertain significance (VUS), 35 as likely pathogenic (LP), and 14 as pathogenic (P). All LP/P variants were located within the conserved olfactomedin domain encoded by exon 3. The updated variant curation guidelines from the Glaucoma VCEP increased the number of clinically definitive classifications from 28% (74/265) to 39% (105/271), with 95% (41/43) of reclassified variants moving to greater clinical relevance. Functional evidence was lacking for 93% (154/166) of VUS. Additional functional evidence could further enhance classification by halving (85/166) the proportion of those classified as VUS. These findings highlight the role of rule calibration and rigorous functional evidence assessment toward improving variant classification with clinical utility for patients.
Hoang, Q. P.; Le, T. X.; Doan, D. D.
Show abstract
Background. Polygenic scores (PRS) for coronary artery disease (CAD) are derived almost entirely from European-ancestry data. Their portability to Southeast Asian populations, including the Vietnamese, is largely uncharacterised and clinically consequential when scores are used with risk thresholds. Methods. We evaluated four independent European-derived CAD scores from the PGS Catalog (PGS000058, PGS000349, PGS002809, PGS004198; 70 - 5,723 variants) in 2,504 individuals from the 1000 Genomes Project, focusing on the Vietnamese Kinh (KHV) and Dai (CDX) samples. Per-individual scores were computed with PLINK2 and standardised. We assessed (i) the cross-ancestry distribution (calibration) and (ii) a clinically-relevant consequence: the proportion of each population flagged high genetic risk when the European top-20% threshold is applied (20% if perfectly calibrated). Results. For the primary score (PGS000058) the standardised PRS differed across super-populations (ANOVA F(4, 2499) = 121.1, p < 0.001); the Vietnamese Kinh mean was +0.47 SD above the European mean (Welch t = 7.77, p = 2.0 x 10^ -14). Applying the European top-20% high-risk threshold, the fraction of Vietnamese Kinh flagged ranged from 22.2% to 57.6% across the four scores, and of Dai from 21.5% to 43.0%, versus the intended 20%. Three of the four scores over-flagged Vietnamese (25-58%); the largest score (PGS004198) was approximately calibrated for East/Southeast Asians ([~]22%) but markedly over-flagged Africans (69.3%). Conclusions. European-derived CAD polygenic scores are inconsistently calibrated in Vietnamese and other Southeast Asian samples, and most substantially over-flag high genetic risk when a European threshold is applied. The magnitude and even the direction of miscalibration depend on the specific score, so no such score can be assumed transferable without local validation and recalibration. Distribution shift bounds, but does not by itself quantify, loss of predictive accuracy, which requires phenotyped data.
Matarage Don, N. N. J.; Biswas, S. B.; Biswas-Fiss, E. E.
Show abstract
Pathogenic mutations in the ABCA4 gene cause several inherited retinal diseases, particularly Stargardt disease (STGD1). However, many missense variants remain classified as variants of uncertain significance (VUS) due to inconclusive evidence regarding their pathogenic impact. The missense VUS span across all the domains of ABCA4, with the majority found in the larger extracellular domains (ECDs). The largest uncharacterized region of ABCA4 is located in ECD1, where limited structural information and inconsistent computational predictions hinder clinical interpretation of missense VUS in this region. Here, we integrated in silico analysis with in vitro functional assays to evaluate the pathogenicity of VUS in this region and improve their diagnostic classification. Missense VUS in the ECD1 uncharacterized region were curated from ClinVar. Six multiallelic sites were identified in the uncharacterized region and 13 missense VUS on these multiallelic sites were characterized using the integrated analysis. In the in silico platform, the pathogenicity of the VUS were predicted using multiple algorithms, and the structural effects of the variants were analyzed compared to the wild type. Recombinant variants were expressed in virus-like particles (VLPs), and protein expression, membrane localization, and ATPase activity were quantified relative to wild type to identify potential disease-causing variants. From the integrated analysis, variants with pronounced structural destabilization, impaired membrane trafficking, and reduced or absent N-retinylidene-phosphatidylethanolamine (NRPE) substrate stimulated ATPase activities were identified as potentially deleterious. Notably, VUS at p.H193P and p.I214N showed loss of function, with p.I214N reflecting selectively impaired membrane targeting and p.H193P reflecting combined expression and trafficking defects. Additionally, NRPE-stimulated ATPase activities were impaired in VUS, p.V195L, p.V195I, p.D197H, p.I214F and p.N269S. Overall structural destabilization interfered with the NRPE-stimulated ATPase activities of p.N269S, while the lack of NRPE-stimulated ATPase activities of p.D197H, p.V195L, p.V195I and p.I214F are thought to be due to impaired NRPE interactions with ABCA4. All the VUS at p.R140, p.H193Y, p.D197N and p.N269H showed both the basal and NRPE-stimulated ATPase activities but less than that of the wild type, displaying a mild functional deficit. Together, these findings demonstrated that certain VUS within the unresolved ECD1 region disrupt ABCA4 stability and function, supporting their contribution to disease pathogenesis. This integrative approach highlights key residues likely to be pathogenic and advances the interpretation of VUS in inherited retinal disorders.
Byun, J.; Saha, D.; Han, Y.; Shaw, V. R.; Siminovitch, K.; Amos, C. I.
Show abstract
BackgroundGenome-wide association studies (GWAS) often fail to identify higher-order epistatic interactions that contribute to complex inheritance patterns of traits and diseases. While machine learning (ML) can capture non-linear relationships, extracting interpretable insights from these models remains a challenge. We propose a novel tree-based feature engineering framework that uses Classification and Regression Trees (CART) to explicitly encode high-order interaction decision paths as dummy variables. We investigate three path-based encoding strategies: (i) all decision paths, (ii) leaf-node paths only, and (iii) internal-node paths only. This approach aims to transform complex decision boundaries into discrete features that capture nonlinear interactions that are not readily captured by traditional association models. ResultsThe framework was evaluated using genetic data for ANCA-associated vasculitis (AAV). To manage the high dimensionality of the engineered feature space, we applied a comprehensive suite of ML methods across three tasks: (1) Ensemble Learning (Random Forest, XGBoost, and Gradient Boosting Machine); (2) Decision Tree Analysis (CART); and (3) Regression and Classification Tasks (Regularized Linear Regression/LASSO, Support Vector Machine, and Logistic Regression). Stepwise feature selection and regularization were employed to isolate the most informative interaction patterns. Results indicate that incorporating CART-derived interaction paths--particularly those from high-impact regions of the tree--significantly improves classification accuracy and model interpretability compared to using the original feature space alone. ConclusionsThe proposed framework provides a robust, scalable methodology for identifying high-order genetic interactions. By bridging the gap between the predictive power of ensemble ML and the necessity for mechanistic insight, this approach offers a clearer mapping of the combinatorial genetic processes underlying complex diseases. While applied here to AAV, the method is highly adaptable for exploring the genetic architecture of diverse populations and complex traits.
Sangkuhl, K.; Whirl-Carrillo, M.; Woon, M.; Venkatesh, R.; Keat, K.; Whaley, R.; Ritchie, M. D.; Klein, T. E.
Show abstract
NAT2 is an important pharmacogene which encodes the N-acetyltransferase 2 enzyme that is involved in the metabolism of multiple medications, and variants in this gene can affect patient response to these medications. CPIC has published a clinical guideline for prescribing hydralazine using NAT2 genotypes. Just prior to the guideline, updated NAT2 star allele numbering and definitions were released, differing somewhat from the historical nomenclature. Clinical pharmacogenomic testing panels often test for the most common star alleles, so knowledge of the most common updated NAT2 star alleles is critical for the implementation of the CPIC NAT2/hydralazine guideline. We first determine NAT2 diplotype frequencies from UK Biobank (UKBB) 200k phased genomes, then analyzed allele, diplotype, and phenotype population frequencies from the All of Us Research program, PennMedicine BioBank (PMBB) and UKBB 500k datasets. We found that analyzing NAT2 diplotypes from phased data provides critical information for algorithms designed to predict diplotypes from unphased data. We observed that NAT2*5, *6, and *4 were the most common star alleles in that order, and the top 11 most frequent NAT2 star alleles were the same across all biobanks. However, differences in star allele frequencies across biogeographical populations were observed. The largest difference led to a higher frequency of NAT2 poor metabolizer phenotypes as compared to rapid and intermediate metabolizer phenotypes in all global populations except in the EAS population, where NAT2 poor metabolizers were in the minority.
Yang, C.; Aguet, F.; Auguste, G.; Ardlie, K. G.; Gerszten, R. E.; Post, W. S.; Wheeler, H. E.; Taylor, K. D.; Kasela, S.; Lappalainen, T.; lm, H. K.; Durda, P.; Johnson, W. C.; Guo, X.; Liu, Y.; Polak, J. F.; Herrington, D. M.; Clish, C. B.; Van Den Berg, D.; Tracy, R. P.; Cornell, E.; Blackwell, T. W.; Papanicolaou, G. J.; Vargas, J. D.; Bekiranov, S.; McNamara, C. A.; Miller, C. L.; Rotter, J. I. I.; Rich, S. S. S.; Manichaikul, A. W.
Show abstract
Introduction: Coronary artery disease (CAD) is a leading cause of death and disability worldwide. Although genome-wide association studies (GWAS) have identified over 300 loci associated with CAD risk, the molecular mechanisms linking these variants to disease and subclinical atherosclerosis are not fully understood. Methods: We performed integration of multi-ancestry CAD GWAS with transcriptomic data from the Multi-Ethnic Study of Atherosclerosis (MESA) obtained through the Trans-Omics for Precision Medicine (TOPMed) program. For integration, we applied Bayesian colocalization analysis with and without statistical fine-mapping to identify genes whose expression levels colocalize with CAD-associated loci. We further applied causal weighted gene co-expression network analysis (cWGCNA) to identify gene co-expression modules and key driver genes associated with subclinical atherosclerosis traits in MESA. Results: We identified 108 genes showing evidence of colocalization with CAD loci, including 24 shared between the two colocalization approaches and 48 novel genes not previously reported in CAD GWAS. Follow-up replication and validation analyses prioritized 5 novel (CCDC30, ZEB1-AS1, ZPR1, PLEKHJ1 and AC018816.3) and 8 previously reported genes (DHDDS, DDX59, LNPEP, DAGLA, ZKSCAN1, LIPA, OPRL1 and EIF2B2) with putative roles in both CAD and subclinical atherosclerosis. cWGCNA identified five gene modules significantly associated with subclinical atherosclerosis in MESA. Additionally, three key driver genes (ATG9B, PRAM1 and ZBTB46) identified by cWGCNA were also identified as CAD-colocalized genes. Discussion: Our integrative analysis highlights key genetic drivers and regulatory networks underlying CAD and subclinical atherosclerosis. These findings underscore the value of incorporating statistical fine-mapping in colocalization studies and demonstrate the utility of combining colocalization with co-expression network analysis to prioritize functional genes and pathways.
Moreau, C.; Morin, G.-P.; Bouchard, J.; Mathieu, J.; Duchesne, E.; Gagnon, C.; Girard, S. L.
Show abstract
Background: Myotonic dystrophy type 1 (DM1) is caused by a CTG repeat expansion in the DMPK gene and represents the most common adult-onset myopathy. Current molecular diagnostics rely on labor-intensive assays that limit accessibility and scalability. Haplotype-based approaches offer a promising alternative for detecting pathogenic expansions indirectly. Methods: We performed genome-wide genotyping in 226 genetically confirmed DM1 patients from the Saguenay-Lac-Saint-Jean founder population and reconstructed haplotypes surrounding the DMPK pathogenic repeat expansion. Based on these haplotypes, we performed a phylogenetic analysis that was further integrated with genealogical reconstruction from the BALSAC database to investigate the origin and transmission of DM1 haplotypes. To evaluate epidemiological utility, we implemented gene dropping simulations within the SLSJ extended genealogies (>80,000 starting individuals) to estimate DM1 incidence at birth. Results: A DM1-associated haplotype was identified in all patients (226/226), consistent with a single major ancestral origin in the SLSJ population. This complete concordance supports the robustness of haplotype-based approaches to infer carrier status without direct repeat sizing. Integrating phylogenetic analysis and genealogical data identified a single couple as the most likely entry point of DM1 in Quebec. Simulation-based estimates of incidence at birth exceeded observed prevalence, suggesting underdiagnosis in the region. Marked geographic heterogeneity in the SLSJ is also observed. Conclusions: Our results demonstrate that haplotype-based approaches can provide a reliable, cost-effective alternative to conventional pathogenic DM1 repeat carriers identification and familial screening strategies.
Jo, J.; Khor, S.-S.; Chu, S.-K.; Ji, Y.; Ueno, K.; Ono, A.; Chen, C.-W.; Do, A.; Han, H.; Kawai, Y.; Kim, N.-E.; Chen, C.-h.; Tokunaga, K.; Won, S.; Yang, H.-C.
Show abstract
Genome-wide association studies (GWASs) have disproportionately focused on European (EUR) populations, limiting the characterization of genetic architecture in other ancestries. To address this imbalance, we integrated large-scale biobanks from Japan, Korea, Taiwan, and China to perform the largest phenome-wide meta-analysis to date in East Asian (EAS) populations, encompassing over one million individuals across 127 complex traits. We identified 8,010 previously unreported associations and observed substantial genetic sharing across EAS subpopulations, while also detecting cohort-specific heterogeneity within the broader EAS context. Transethnic analyses revealed moderate genetic correlations between EAS and EUR populations, indicating both shared and ancestry-specific components of disease risk. Pleiotropy analyses highlighted prominent signals within the HLA region, supported by protein-protein interaction connectivity and immune-related pathway enrichment. Decomposition of genome-wide association matrices further uncovered structured cross-trait architectures, revealing a predominantly shared polygenic backbone driven by metabolic, biochemical, and anthropometric traits, together with two discrete latent components enriched for immune-related processes. Together, our findings refine the genetic architecture of complex traits in East Asian populations at unprecedented scale and clarify the balance between shared and population-specific determinants of human diseases.
NESHATUL, H.; Wagenknecht, J.; Dong, X.; Zimmermann, M. T.
Show abstract
Evaluating the impact of genomic variation is essential for identifying underlying mechanistic causes of human diseases. The spectrum of neurodevelopmental disorders is driven by diverse genetic alterations with genes like SMARCA4 being prototypical examples. There have been significant hurdles to implementing the protein-specific and mechanism-informed variation effect predictors that are anticipated to have the highest yield of mechanistic information. Yet, there is a pressing need, for example, within SMARCA4 where 98% of the 2780 reported variants lack a disposition and remain of uncertain significance (VUS). Further, the field has yet to identify each variants specific molecular mechanism, which will inform targeted therapeutic development strategies. In this study we developed a mechanistic structure-informed helicase-specific variant effect predictor by leveraging diverse information with state-specific calculations. Our approach has 100% recall of pathogenic variants while classifying 87.23% of VUS into damaging (55.74%, n=262) versus tolerated effects (31.49%, n=148), including those with conflicting interpretations. This analysis reveals significant enrichment of integrated functional metrics, such as conservation and solvent exposure, that parallel allele frequences in health populations, and emphasizes the robustness of the method. Thus, we have demonstrated a novel approach for the development of mechanism-informed protein-specific interpretation of human genetic information.
Rodenburg, K.; Fenwick, L.; Pennings, R.; Haer-Wigman, L.; Ben-Yosef, T.; van Erp, F.; Reurink, J.; Gilissen, C.; van den Born, L. I.; Cremers, F. P. M.; Cohen, Y.; Yntema, H.; de Vrieze, E.; Kremer, H.; de Bruijn, S. E.; Collin, R. W. J.; Roosing, S.; van Wijk, E.
Show abstract
Despite substantial advances in diagnostic testing, 10-15% of Usher syndrome patients remain without a genetic diagnosis, having significant implications for genetic counseling and potential future therapeutic interventions. In this study, genome sequencing data from probands clinically presenting with Usher syndrome were analyzed. Two novel deep-intronic variants were identified in PCDH15, c.3983+3635A>G and c.3123-1728A>G, in two independent patients. Both deep-intronic variants were classified as likely pathogenic and predicted to alter PCDH15 pre-mRNA splicing. Using a minigene splice assay and iPSC-derived photoreceptor precursor cells from patients, we confirmed that both variants lead to the inclusion of a pseudoexon in the PCDH15 transcript introducing a stop codon and subsequent premature termination of protein translation. We designed and evaluated antisense oligonucleotides (ASOs) with the purpose of redirecting aberrant pre-mRNA splicing caused by both deep-intronic variants. For both variants, designed ASOs were successful in restoring normal splicing patterns, highlighting their potential as a future therapeutic intervention strategy to halt the progression of retinitis pigmentosa caused by these novel variants. Overall, these findings contribute to the understanding of Usher syndrome caused by deep-intronic pathogenic variants in PCDH15 and describe for the first time the use of an ASO-mediated splice correction strategy for individuals diagnosed with these variants.
Uren, C.; Moller, M.; Oelofse, C. R.
Show abstract
Tuberculosis (TB) remains a major public health challenge, exerting profound socio-economic burdens and causing debilitating illness in approximately 2.5 million individuals across Africa annually. Optimized large-scale treatment regimens, such as NAT2-genotype adjusted dosing, could improve patient outcomes and strengthen healthcare systems. However, fully addressing the complexity of multi-drug TB treatment responses requires consideration of the entire pharmacogenomic (PGx) landscape, particularly within African populations, which are both genetically diverse and critically understudied. In this study, we predict NAT2 genotypes and phenotypes in specific African populations, and we extend TB PGx research beyond well-established biomarkers. Current bioinformatic prediction tools were used to evaluate individual- and population-specific variation in genotype and next-generation sequencing data from 2,143 individuals across 20 African population groups, spanning ten PGx genes associated with multi-drug TB treatment and response. Most predicted functionally deleterious variants occurred at low frequencies (MAF < 0.01) and were observed in only one of the 20 populations. The Khomani and Nama populations had a distinctly higher proportion of NAT2 fast metabolizer phenotypes than other African populations, indicating a lower risk of INH overexposure and possibly different dosage requirements in these groups. These findings highlight both the potential and current limitations of functional prediction for absorption, distribution, metabolism and excretion (ADME) variants, and the transferability of their predictive value between African population groups. With the increasing accessibility of next-generation sequencing, alongside the development of comprehensive databases capturing African variation and advances in computational algorithms, the cumulative impact of genetic variation on TB drug response can be more accurately captured, thereby informing precision treatment strategies.
Yang, Q.; Zou, W.-B.; Pu, N.; Li, Y.; Hu, Y.; Wang, Y.-C.; Liu, X.; Genin, E.; Masson, E.; Wang, J.; Ferec, C.; Cooper, D. N.; Li, W.; Chen, J.-M.
Show abstract
As genomic sequencing evolves beyond rare disease diagnostics toward population screening and precision medicine, clinical variant interpretation is increasingly challenged by variants whose clinical consequences depend on biological context. Current frameworks, including the ACMG/AMP guidelines, generally assign a single classification to each variant regardless of inheritance state or genetic context, potentially failing to communicate context-dependent clinical consequences. Here, we address this issue using loss-of-function variants in LPL as a uniquely informative model system in which residual physiological LPL activity can be directly quantified in vivo. By systematically integrating published biallelic LPL genotypes, physiological measurements, functional studies, and clinical phenotypes, we identified a biologically meaningful transition at approximately 10% residual physiological LPL activity. Activity below this level was predominantly associated with classical childhood-onset familial chylomicronemia syndrome (FCS), whereas higher activity was associated with phenotypic attenuation and modifier-dependent clinical expression. Furthermore, heterozygous loss-of-function variants exhibited an estimated penetrance of 5-7% for severe hypertriglyceridemia. We therefore propose a context-dependent framework in which biallelic complete- or near-complete loss-of-function genotypes are interpreted as causative for FCS, whereas heterozygous variants are interpreted as predisposing to severe hypertriglyceridemia while retaining recognition of FCS carrier status. Together, our findings demonstrate that clinical variant interpretation should integrate available biological context--including, where relevant, allelic configuration, residual biological function, and penetrance--rather than rely on the intrinsic molecular consequence of the variant alone. More broadly, this framework provides a conceptual model for interpreting variants across the continuum from Mendelian disease to genetic predisposition in the era of precision medicine.
Happ, H.; Christensen, B.; Knight, S.; Novoa, A.; Isakson, D.; Nadauld, L.; Quinlan, A.; Bonkowsky, J. L.
Show abstract
Background and Objectives: Leukodystrophies are rare genetic diseases affecting the central nervous system white matter, leading to progressive disabilities and death. Although early diagnosis is critical for therapies, the penetrance and phenotypic spectrum of many leukodystrophies remain poorly defined. Here, we integrate sequencing population screening with longitudinal electronic health record (EHR) data. Our goals were to assess the prevalence of undiagnosed leukodystrophy, characterize phenotypic variability among genotype-positive individuals, and estimate penetrance across multiple leukodystrophies. Methods: We analyzed 19 genes associated with 13 leukodystrophies in pediatric and adult individuals recruited via the HerediGene Population Study, a 5-year study conducted primarily of healthy individuals in the U.S. intermountain west. Sequencing was performed on 210,983 individuals, consisting of genome sequencing for 34,033 and SNP panel imputation for 176,950. Variant results were cross-referenced to comprehensive and longitudinal (20+ years) clinical data in the Intermountain Health Enterprise Data Warehouse and to the Utah Leukodystrophy Program. Results: Pathogenic variants were identified in 4 genes (CSF1R, PLP1, POLR3A, SNORD118) in 9 individuals, none of whom had a clinical leukodystrophy diagnosis or characteristic MRI findings. These findings suggest that missed clinical diagnoses of most leukodystrophies are uncommon in a centralized healthcare system, but also demonstrate that for some leukodystrophies there may be variable or reduced penetrance, or broader phenotypic spectra than recognized. We used published incidence estimates and the observed leukodystrophy-associated genotypes to infer penetrance ranges that varied from wide for ultra-rare leukodystrophies, to tightly bounded for more prevalent conditions. Discussion: In this predominantly healthy population, we did not find any patients with leukodystrophy who had been genetically undiagnosed but then identified by sequencing. However, we identified 9 individuals with genotypes previously reported to result in leukodystrophy, but none of whom had clinical symptoms or MRI features associated with the specific leukodystrophy. Our results support a revised model in which leukodystrophies exist along a continuum of penetrance and expressivity, with implications for newborn screening, variant interpretation, and risk stratification.
Kandasamy, R.; Gurung, M.; Shrestha, S.; Bibi, S.; Thorson, S.; Carter, M.; O'Connor, D.; Murdoch, D. R.; Kelly, D. F.; Shrestha, S.; Levin, M.; Pollard, A. J.
Show abstract
Background Pneumococcal disease is a leading cause of paediatric pneumonia and meningitis. Pneumococcal colonisation is the fundamental step to pneumococcal disease causation. We aimed to identify genetic loci associated with pneumococcal colonisation amongst children. Methods We conducted a genome-wide association study on 2111 Nepalese children, comprising 1346 cases carrying pneumococcus and 765 controls. We tested 8.1 million imputed variants using logistic regression and ten principal components as covariates. Fine mapping and functional evidence were used to identify suspected causal variants and related genes of interest. Findings A cluster of 22 variants of genome-wide significance (p<5x10-8) were identified on chromosome 12q21.31, eight of which were within PPFIA2. Fine mapping of this region identified 5 variants within 0.1 Mb of the 5-prime region of PPFIA2 all of which are significant eQTLs for PPFIA2. We further describe three loci (10q23.31, 12q23.1, and 20p11.21) which had variants with highly suggestive associations (p<5x10-7)with pneumococcal carriage. Interpretation Our study demonstrate human susceptibility to pneumococcal carriage to be polygenic with genetic variations which regulate PPFIA2 expression playing a key role in the ability for pneumococcus to colonise children. Targeting these genetic factors and the associated pathways are a means for preventing pneumococcal disease. Funding This study was supported by funding from Gavi - the vaccine alliance, the European Unions Horizon 2020 research and innovation program under grant agreement number 668303 (PERFORM), and a Robert Austrian Research Award.
Sankaranarayanan, R.; Vasavada, A. R.; Agrawal, D.; Vasavada, S. A.; Vasavada, V. A.
Show abstract
Purpose: To identify transcript-level variants in crystallin genes in paediatric patients with unilateral cataracts. Methods: Anterior capsulorhexis (n=12) from patients underwent surgical management of congenital unilateral cataracts was collected. Total RNA was isolated from lens epithelial cells, and complementary DNA (cDNA) was synthesized. Full-length RNA transcripts of 10 lens-specific crystallin genes were PCR-amplified and analysed via Sanger sequencing. Identified transcript variants were further validated using genomic DNA (gDNA) through Sanger sequencing. In addition, the full-length (~7,535 bp) CRYBA1 genomic region was sequenced using Oxford Nanopore Technology. Results: Aberrant low molecular weight (LMW) amplicons (~370 bp) of the CRYBA1 transcript were identified in three patients presented with unilateral cataract. Of 3 patients, 2 had persistent fetal vasculature (PFV) and 1 had pre-existing posterior capsular defect (PPCD). Sanger sequencing revealed a precise loss of exons 2 to 4 in the CRYBA1 RNA transcript. No coding, splice-site, or large deletion variants were detected in the genomic DNA of the patients or their parents. In silico analysis predicted two possible truncated proteins arising from these alternatively spliced transcripts: one comprising the first 11 amino acids of the N-terminal region with a loss of all Greek key motifs, and another comprising 90 amino acids encoded by exons 5 and 6, initiated from an alternative start codon in exon 5, and loss of Greek key motifs 1 & 2. Conclusion: The precise skipping of exons 2 to 4, consistent with canonical splicing signals (5-prime-GU...AG-3-prime), in the absence of genomic alterations, suggests the presence of alternatively spliced (AS) CRYBA1 transcripts in human lenses. This is the first report documenting AS-CRYBA1 transcripts in association with childhood cataracts with PFV and PPCD.
Chen, S.; Moorthy, A.; Yu, P. K.; Wang, J.; Liu, D.
Show abstract
With the increasing accessibility of single-cell RNA sequencing (scRNA-seq) data, cell-type-specific gene expression can be linked to complex traits through pseudo-bulk method, which considered aggregated gene expression from multiple cells of the same annotated cell type per individual and clearly shows the limitation of ignoring intra-individual cell-to-cell variability. Concurrently, pseudotime trajectory inference has gained popularity for its ability to capture continuous biological processes such as cell differentiation and lineage development, instead of individual discrete stages. It is natural to consider whether genetic effects for complex traits, such as individual level disease status, show a dynamic pattern along the inferred trajectories. In this study, we introduce a novel framework that models gene expression as a function of pseudotime along the inferred trajectories. We mapped expression quantitative trait loci (eQTL) effects in the cis-region as functional parameters, which we called "dynamic eQTLs", showing regulatory effects exerted by genetic variants change continuously along the cellular trajectory. For eQTLs of constant effects across pseudotime we leveraged external bulk-eQTL information to enhance the power. Furthermore, we employed significant, variable dynamic eQTLs as instrumental variables to infer causal relationships between gene expression and complex traits. To address challenges inherent to scRNA-seq data--such as sparsity and high variability--we incorporate an empirical likelihood-based inference method, which is non-parametric and self-normalized. Besides, genes associated with trajectory branchpoints may bring confounding, and we also proposed a causal mediation analysis framework to determine whether a gene plays a causal role for the disease directly and indirectly through driving cell fates. Applying our method to scRNA-seq data from human lung tissue of 114 samples (66 pulmonary fibrosis cases and 48 controls), along with meta-analyzed GWAS summary statistics for IPF from 3 studies, we identified pseudotime-dependent causal effects for IPF from genes implicated in the trajectory AT2 - translational AT2 - AT1, which is crucial in lung tissue repair and regeneration. We also found that 30 genes have a mediated effect through cell fates.
Dibbasey, M.; Esoh, K.; Susso, B.; Forrest, K.; Sonko, B.; Makalo, L.; Oriero, E.; Cheng, N. I.; Amenga-Etego, L.; Cerami, C.; Amambua-Ngwa, A.
Show abstract
Globally, approximately 75% of sickle cell disease (SCD) cases occur in sub-Saharan Africa, yet empirical data on its natural history, clinical burden, and modifiers remain scarce in the region. This retrospective study describes the demographic characteristics, complications, and routine care and examines how non-genetic factors and blood markers relate to disease severity. We analysed 8402 medical records from 840 SCD patients with confirmed HbSS genotype registered in MRCG Keneba and Fajara clinics (NKeneba=148; NFajara=692). A generalised linear model was employed to estimate the association of non-genetic correlates, blood biomarkers, and routine care medications with disease severity. Here, we showed 67% of patients in the Keneba cohort and 92% of those in the Fajara cohort had no documented SCD-related chronic complication. Despite no documented evidence of hydroxyurea use, rates of SCD crises (Keneba=0.57, Fajara=0.63) and infections (Keneba=0.53, Fajara=0.35), expressed per patient-year, were low in both cohorts, with 99% of patients experiencing less than or equal to 3 SCD crises per patient-year. Age at diagnosis, gender and seasonality were not significantly associated with SCD crises or other clinical outcomes/events rates. Each additional folic acid prescription was associated with higher haemoglobin(g/dL) (total folic acid prescriptions: Beta-Fajara=1.31, P=0.005; Beta-Keneba=1.20, P<0.001). Penicillin prophylaxis was associated with a reduced rate of infection (total Pen V prescriptions: IRR-Fajara=0.85, P=0.002; IRR-Keneba=0.93, P=0.002) and SCD crises (IRR-Fajara=0.67, P=0.001; IRR-Keneba=0.87, P=0.001). This study found low acute event rates and chronic complications prevalence in the absence of hydroxyurea use. No significant associations were observed between non-genetic correlates and clinical events, but the study highlighted the need for continued folic acid supplementation and penicillin prophylaxis due to their observed beneficial effects.